java 抓取网页内容实现代码
package test;import java.io.BufferedReader;
import java.io.IOException;
import java.io.InputStream;
import java.io.InputStreamReader;
import java.net.Authenticator;
import java.net.HttpURLConnection;
import java.net.PasswordAuthentication;
import java.net.URL;
import java.net.URLConnection;
import java.util.Properties;
public class URLTest {
// 一个public方法,返回字符串,错误则返回"error open url"
public static String getContent(String strUrl) {
try {
URL url = new URL(strUrl);
BufferedReader br = new BufferedReader(new InputStreamReader(url
.openStream()));
String s = "";
StringBuffer sb = new StringBuffer("");
while ((s = br.readLine()) != null) {
sb.append(s + "/r/n");
}
br.close();
return sb.toString();
} catch (Exception e) {
return "error open url:" + strUrl;
}
}
public static void initProxy(String host, int port, final String username,
final String password) {
Authenticator.setDefault(new Authenticator() {
protected PasswordAuthentication getPasswordAuthentication() {
return new PasswordAuthentication(username,
new String(password).toCharArray());
}
});
System.setProperty("http.proxyType", "4");
System.setProperty("http.proxyPort", Integer.toString(port));
System.setProperty("http.proxyHost", host);
System.setProperty("http.proxySet", "true");
}
public static void main(String[] args) throws IOException {
String url = "//www.jb51.net";
String proxy = "http://192.168.22.81";
int port = 80;
String username = "username";
String password = "password";
String curLine = "";
String content = "";
URL server = new URL(url);
initProxy(proxy, port, username, password);
HttpURLConnection connection = (HttpURLConnection) server
.openConnection();
connection.connect();
InputStream is = connection.getInputStream();
BufferedReader reader = new BufferedReader(new
InputStreamReader(is));
while ((curLine = reader.readLine()) != null) {
content = content + curLine+ "/r/n";
}
System.out.println("content= " + content);
is.close();
System.out.println(getContent(url));
}
}
您可能感兴趣的文章:
- java利用url实现网页内容的抓取
- java利用url实现网页内容的抓取
- 简单的java爬虫抓取网页实现代码(未测试)
- 网络爬虫Java实现抓取网页内容
- java代码抓取网页邮箱的实现方法
- java利用url实现网页内容的抓取
- java利用url实现网页内容的抓取
- java利用url实现网页内容的抓取
- 爬虫技术(2)--抓取网页java代码实现
- php 实现信息采集(网页内容抓取)程序代码
- java利用url实现网页内容的抓取
- java 抓取 https 网页内容
- js 实现打印网页中定义的部分内容的代码
- jQuery 处理网页内容的实现代码
- PHP 抓取网页图片并且另存为的实现代码
- Java 抓取网页内容,获取指定服务器IP
- select 控制网页内容隐藏于显示的实现代码
- PHP 抓取网页图片并且另存为的实现代码
- 【JAVA】 抓取网页内容