The idea
A web page is a set of files, and each file needs fetching. The question is how many TCP connections that takes and how far each request has to travel. Open a fresh connection for every picture and you pay set-up cost every time. Make every request travel to a distant server and a slow link becomes the bottleneck.
HTTP answers the first problem with persistent connections. A cache near the client answers the second by keeping copies of what was recently fetched.
How it works
The Web and URLs
The World Wide Web (WWW) is an information space where documents and other web resources are identified by Uniform Resource Locators (URLs), interlinked by hypertext links and accessible via the Internet. Today it is a distributed client-server service. The web client is a browser (the lecture lists Internet Explorer, Firefox, Chrome and Safari). The web server is where the page is stored (Apache and Microsoft Internet Information Server).
A URL is the address of a resource on the Internet. It gives the location of the resource and the protocol to reach it. It usually has four parts:
- Protocol: the program needed to access the page, for example
httporftp. - Host: the IP address of the server, or its domain name.
- Port: a 16-bit port number, which can be omitted if it is well known.
- Path: the location and name of the file in the underlying operating system.
Different separators sit between the four parts. The lecture’s example is
https://sydney.edu.au/engineering/about/school-of-electrical-and-information-engineering.html.
How it works
HTTP
HyperText Transfer Protocol (HTTP) defines how client and server programs
are written to retrieve web pages. It is the web’s application layer
protocol. The server uses port 80 and HTTP uses the services of TCP.
How it works
Non-persistent and persistent connections
With non-persistent connections each object, such as a picture, is retrieved over a new TCP connection. One TCP connection is made for each request and response. This is how HTTP worked before version 1.1, and the overhead is high because every request needs its own connection.
With persistent connections one TCP connection retrieves all the objects. The server leaves the connection open for more requests after sending a response. It closes the connection when the client asks it to, or when a time-out is reached. HTTP version 1.1 specifies a persistent connection by default.
Aside
The lecture's HTTP example is missing here
The lecture shows one diagram for the non-persistent case, captioned “connections are opened, closed, opened and closed”, and one for the persistent case, captioned “connections are always open”. Both are figures and nothing from them extracted except those captions. This page gives no round-trip count or timing for either case, because the lecture’s numbers are not available. Check the slides if a question asks for response time in round-trip times.
How it works
Web caching with a proxy server
A proxy server is a computer that keeps copies of responses to recent requests. The HTTP client sends its request to the proxy. If the object is in the cache, the cache returns it. Otherwise the cache requests the object from the origin server, then returns it to the client.
The lecture’s illustration has client A ask for a video at time T. The proxy
fetches it from the origin server and keeps a local copy. When client B asks
for the same video at time T + t, the proxy serves its copy.
The advantages are that it reduces the load on the original server, decreases traffic and improves latency.
The lecture assumes the cache is close to the clients, for example in the same network. Response time is smaller because the cache is closer, and traffic drops on the link out of the institutional or local ISP network, which is often the bottleneck.
- L
- Object (packet) length in bits
- R
- Link bandwidth in bps
- a
- Average arrival rate in requests per second
- rho
- Traffic intensity (utilisation), dimensionless
- D
- Average delay in seconds, including queueing
How it works
Queueing delay and traffic intensity
The lecture models queueing delay with the M/M/1 model. The traffic intensity
is La/R, written as the channel occupancy. When La/R is close to 0 the
average queueing delay is small. As La/R approaches 1 the delay becomes
large. When La/R is greater than 1 more work arrives than can be
serviced, and the average delay is infinite. The effective transmission rate
is (1 - rho)R bps.
This is why a link at 0.97 utilisation is a problem, even though it is
under 1.
Worked example
Caching example, original 1.54 Mbps access link
The scenario: access link rate 1.54 Mbps, round trip from the public
Internet router to the origin server 2 s, object size 100 K bits,
average request rate 15 per second, LAN 1 Gbps.
- Data rate to browsers:
15 requests/s x 100 kbit = 1.5 Mbps. - LAN utilisation:
1.5 Mbps / 1 Gbps = 0.0015. - Access link utilisation:
1.5 Mbps / 1.54 Mbps = 0.97. - Access link delay:
100 kbit / (1.54 Mbps x (1 - 0.97)) = 2.165 s. - LAN delay:
(100 kbit / 1 Gbps) / (1 - 0.0015) = 0.1 ms. - End-to-end delay:
2 s + (2 x 2.165) s + (2 x 0.0001) s, about6.3 s.
AnswerEnd-to-end delay about 6.3 s
Worked example
Caching example, buy a 154 Mbps access link
The same scenario with the access link raised to 154 Mbps.
- Access link utilisation:
1.5 Mbps / 154 Mbps = 0.0097. - Access link delay:
100 kbit / (154 Mbps x (1 - 0.0097)) = 0.66 ms. - The delay on that link drops from seconds to milliseconds, so the 2 s round trip now dominates.
- The lecture’s cost note: a faster access link is expensive.
AnswerAccess link delay about 0.66 ms, cost is high
Worked example
Caching example, install a web cache with a 0.4 hit rate
The access link stays at 1.54 Mbps and a cache in the institutional network
satisfies 40% of requests.
- Requests that use the access link:
60%. - Data rate over the access link:
0.6 x 1.50 Mbps = 0.9 Mbps. - Access link utilisation:
0.9 / 1.54 = 0.58. - Delay for a request served from the origin:
RTT 2 s + 0.32 s + 2 x 0.0001 s, where0.32 sis the lecture’s figure for the access link term. - Delay for a request served from the cache: about
2 x 0.0001 s. - Average:
0.6 x (2 + 0.32 + 0.0002) + 0.4 x (0.0002), about1.392 s.
The lecture notes this is a lower average delay than the 154 Mbps link, and cheaper too.
AnswerAverage delay about 1.392 s, cheaper than the 154 Mbps link
Aside
A small mismatch in the 0.32 s figure
Applying the delay formula to the cached utilisation of 0.58 gives
100 kbit / (1.54 Mbps x 0.42), about 0.155 s, and twice that is about
0.31 s. The slide prints 0.32 s, and its total of 1.392 s follows from
0.32 s. Follow the slide’s value in an exam. The difference is rounding in
the utilisation figure and does not change the conclusion.
Where marks get lost
The 2 s in the original example is the round trip from the public Internet
router to the origin server. It is not reduced by a faster access link, so
buying 154 Mbps removes the queueing delay but leaves about 2 s of delay
in place. The cache beats it because cached requests skip that round trip
entirely.
In the exam
- URL parts: protocol, host, port, path. The port is 16 bits and can be omitted when well known.
- HTTP: server port
80, runs over TCP. Persistent is the default from version 1.1. Non-persistent uses one TCP connection per object. - Proxy advantages: less load on the origin server, less traffic, better latency.
- Caching calculation: be ready to give the access link utilisation
(
0.97), the delay with the original link (about6.3 s), with the faster link (0.66 mson the link) and with the cache (0.9 Mbps, utilisation0.58, about1.392 s). - Queueing: delay is small for traffic intensity near
0, large near1, infinite above1.
Check yourself
- A URL has protocol, host, port and path. HTTP runs on TCP port
80. - Persistent connections reuse one TCP connection. HTTP 1.1 makes them the default.
- A proxy serves repeat requests from its own copies and so cuts origin load, traffic and latency.
- At
0.97utilisation the original link gives about6.3 s. A0.4hit rate cache brings the average to about1.392 s.